Biology Methods and Protocols
◐ Oxford University Press (OUP)
Preprints posted in the last 7 days, ranked by how well they match Biology Methods and Protocols's content profile, based on 61 papers previously published here. The average preprint has a 0.08% match score for this journal, so anything above that is already an above-average fit.
Chaturvedi, R. R.; Gracner, T.; Perez-Arce, F.; Suen, S.-c.; Jin, J.; Orriens, B.; Pacula, R. L.; Sexton Ward, A.; Haile, R.; Kapteyn, A.
Show abstract
Importance: Evidence on GLP-1/GIP therapies is largely derived from trials enrolling selected populations or medical records that miss utilization outside healthcare channels. No nationally representative cohort has characterized real-world uptake, indications, and access. Objective: To characterize GLP-1/GIP prevalence, indication, clinical profile, and access. Design: Prospective cohort study with three GLP-1/GIP surveillance waves (March 2024, December 2024, October 2025). Setting: The Understanding America Study, an address-based, nationally representative panel of approximately 15,000 US adults aged 18+ years initiated in 2014. Participants: UAS participants responding to at least one surveillance wave (n=9150). Exposures: GLP-1/GIP use status (never vs any use, comprising current and former use), self-reported primary indication (diabetes, weight loss, or other), and access pathway (traditional vs non-traditional). Main Outcomes and Measures: Survey-weighted prevalence of GLP-1/GIP use, overall and by indication and access pathway; sociodemographic, cardiometabolic, treatment, and access characteristics; and smartwatch-derived resting heart rate, heart rate variability, maximum activity heart rate, step count, and sleep duration and variability. Results: Among n=9150 adults (1274 with any use; 60.9% female; median age 53 years), weighted prevalence increased 46%, from 8.2% (March 2024) to 12.0% (October 2025) representing 32 million. Weight-loss indications grew, reaching nearly half of use (4.1% to 5.6%); diabetes-indicated use was stable (5.3% to 5.4%). Users carried high cardiometabolic burden (obesity, 68.2%; diabetes, 53.6%) but diverged by indication: diabetes-indicated users were older (median, 59 vs 49 years), whereas weight-loss-indicated users were more often female (69.9% vs 51.3%) and healthier. One in three users (~9 million) had non-traditional access, especially in weight-loss-indicated users, of whom 33% had no conventional prescription; 41% used compounding, online, or foreign pharmacies; and, 43% lacked coverage. Non-traditional users were five times as likely to report an unlisted, likely compounded formulation (19.8% vs 4.1%). All p<0.05. Conclusions and Relevance: Real-world GLP-1/GIP use has grown rapidly and diversified substantially in indication, access, and population profile. One in 3 users obtained treatment through nontraditional channels largely invisible to claims data, raising long-term safety, efficacy, and coverage questions. GLIMMER provides a public, nationally representative longitudinal evidence base for future payer and provider decisions.
Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.
Show abstract
Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
Hessel, M.; Inda Diaz, J. S.; Sjöberg, A.; Salva-Serra, F.; Helldal, L.; Jirstrand, M.; Johnning, A.; Kristiansson, E.; Skovbjerg, S.
Show abstract
Antimicrobial resistance is a public health challenge, driving the need for rapid, cost-effective diagnostic support tools. Artificial intelligence (AI) may enable prediction of susceptibility to untested antibiotics from known susceptibility results, but prospective clinical validation is required before routine use. We evaluated an AI-based decision support method, trained on invasive isolates from the European Surveillance System (TESSy), for prediction of antibiotic susceptibility in clinical Escherichia coli urine isolates. The evaluation included 99 E. coli isolates from urine samples with diversity in age, sex, and antibiotic susceptibility. Predictions were evaluated for 14 antibiotics using patient metadata and susceptibility results for 4-8 antibiotics as input. Prediction uncertainty was handled using conformal prediction, allowing abstention when confidence was insufficient. EUCAST disk diffusion test results were used as reference and genomic sequence data was used to explore mechanisms of the AI performance. Without conformal prediction, 84% of predictions were correct when susceptibility results of six antibiotics were used to predict susceptibility to eight additional antibiotics. Across all predictions generated using susceptibility results for six antibiotics as input, the major and very major error rates were 19% and 12%, respectively. Prediction errors varied between antibiotics and were associated with certain phenotypic and genotypic resistance patterns. Conformal prediction reduced errors but increased abstentions; at confidence levels of 90%, 95%, and 97.5%, the model abstained in 9.6%, 14%, and 22% of instances. The method showed promising performance, but its clinical use remains limited and may require diagnostic data beyond susceptibility test results and demographic variables.
Tiwari, P.; Garg, M.; Pattanayak, S.; Sarkar, I.; Roy, R.; Bhatraju, N.; Verma, A.; K, S. R.; Prakash, S.; Kumar, V. S.; Uddin, M. A.; Rawat, N.; Sahu, A.; Kumar, Y.; Leuva, P. H.; Mridha, A.; Yenamandra, V.; Singh, A. P.; Mishra, A.; Raychaudhuri, S.; Tallapaka, K. B.; Chandak, G. R.; Kulkarni, M. J.; Dharne, M.; Wahengbam, R.; Kalita, J.; Manna, P.; Subudhi, U.; Majumder, S.; Chakraborty, P.; Chaudhary, K.; Sengupta, S.; Phenome India Consortium, ; Sardana, V.; Chatterjee, S.; Ganguly, D.
Show abstract
Background: India has a rising incidence of chronic non-communicable diseases, making it a major healthcare burden today. Growing evidence suggests that chronic low-grade inflammation links ageing with cardiometabolic disorders, captured by the emerging concept of inflammaging. However, most evidence on biological ageing comes from Western populations, with no similar models developed for the Indian population. Given the country's distinctive genetic makeup, unique exposome, and heterogeneous NCD presentation, Western models may not capture inflammaging and its effects in the Indian population. Methods: We analysed baseline data from 4,240 adults in the Phenome India CSIR Health Cohort Knowledgebase (PI CheCK), a nationwide multi-centre cohort. Participants were stratified into eight cardiometabolic phenotype groups by BMI (Asian cut off), blood pressure and HbA1c status. We trained a Super Learner ensemble to predict chronological age in the lean normotensive-normoglycaemic reference group (n=615) using 44 plasma cytokines, sex, haemoglobin, and bioimpedance-derived visceral fat area, per cent body fat, and total body water. Performance was assessed by repeated five-fold cross-validation and in a held-out healthy test set. Calibrated biological age acceleration was then estimated in the remaining 3,625 participants. Results: Median age was 51.0 years (IQR 41.0 to 62.0) and 49.4% were female. The Super Learner outperformed elastic net and XGBoost comparators. Permutation importance identified visceral fat area, per cent body fat, CTACK, SDF1a, haemoglobin and sex as leading contributors, with body composition measures accounting for the largest share, indicating an immune-metabolic rather than cytokine-only signal. Biological age acceleration was concentrated in overweight/obese phenotypes. Lean phenotypes showed acceleration close to the reference (0.32 0.50 years). Conclusions: Cytokine and body composition measures capture a quantifiable immunometabolic ageing signal in a South Asian cohort, with acceleration driven predominantly by adiposity. External validation and longitudinal follow up are required.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Corzantes, K.; Choy, K.; Adar, S.; Castellanos, L. F.; Gross, A. L.; Langa, K. M.; Rohloff, P.; Weerman, B.; Briceno, E.; Ramirez-Zea, M.; Behrman, J.; Flood, D.
Show abstract
Introduction Guatemala is the most populous country in Central America and a setting with unique opportunities for aging research. Approximately 40% of Guatemala's population is Indigenous Maya, who together speak 22 Mayan languages. Currently, there is no population-based aging study in Guatemala and few aging studies in Latin America among Indigenous populations. The Longitudinal Study of Aging in Guatemala (ELEGUA) aims to address these gaps by developing a nationally representative, population-based, longitudinal aging study modeled on the Health and Retirement Study and the Harmonized Cognitive Assessment Protocol, adapted to the cultural and linguistic context of Guatemala. The objective of this protocol is to describe the rationale and design of the ELEGUA pilot survey. Methods and analysis The ELEGUA pilot was a cross-sectional household survey of adults aged 40 years or older in Tecpan, Guatemala. Tecpan was chosen because its diverse population facilitated testing of study procedures in both Spanish and Kaqchikel, a common Mayan language. The survey included up to 600 households sampled using a multistage stratified cluster design. Within each household, one individual aged 40 years or older was selected, with oversampling of adults aged 55 years or older. This respondent completed a comprehensive questionnaire, including detailed cognitive tests, and provided physical measurements and a venous blood sample. Household respondents provided information on household economics and family structure, and an informant reported on the individual respondent's cognitive function. Data were collected using a computer-assisted personal interviewing system. Planned analyses include survey-weighted descriptive statistics and psychometric evaluation of the cognitive assessments. Ethics and dissemination Ethics approval was obtained from the ethics committees of the Institute of Nutrition of Central America and Panama, Maya Health Alliance, and the University of Michigan. Results will be disseminated through publications in peer-reviewed journals and presentations to local, national, and international audiences.
Green, J. L.; Davies, H.; Russell, D. A.
Show abstract
Background: The relative merits of infrainguinal bypass and primary major lower limb amputation (MLLA) for chronic limb-threatening ischaemia (CLTI) remain uncertain, and the baseline profiles of patients selected for each strategy are poorly described. Methods: A systematic review and meta-analysis were undertaken in accordance with PRISMA 2020 and prospectively registered (PROSPERO: CRD42022356094). MEDLINE, Embase, CENTRAL, and CINAHL were searched from inception to March 2025. Prospective studies of adults with CLTI undergoing primary infrainguinal bypass or primary MLLA were eligible. Mortality, major adverse cardiovascular events (MACE) and subsequent amputation outcomes were synthesised using random-effects meta-analysis of proportions. Baseline comorbidity profiles were also extracted. Results: Twenty-seven studies involving 6,576 patients were included: 5,779 underwent infrainguinal bypass and 797 underwent MLLA. After bypass, pooled mortality was 3.7% at 30 days (95% CI 2.8%-4.9%, I2 = 49.4%), 18.5% at 1 year (95% CI 15.6%-21.9%, I2 = 62.3%), and 54.3% at 5 years (95% CI 50.5%-58.0%, I2 = 0%). After MLLA, pooled mortality was 9.2% at 30 days (95% CI 4.1%-19.3%, I2 = 73.5%), 28.5% at 1 year (95% CI 13.3%-51.0, I2 = 70.8%), and 39.9% at 2 years (95% CI 0.3%-99.3, I2 = 90.5%), although longer-term estimates were limited by sparse data and marked heterogeneity. Thirty-day MACE was 6.5% (95% CI 4.3%-9.7, I2 = 63.5%) after bypass and 2.8% after MLLA (95% CI 0.1%-37.6%, I2 = 0%). Early subsequent major amputation after bypass occurred in 3.9% of patients (95% CI 2.0%-7.7%, I2 = 91.2%), rising to 16.2% at 1 year (95% CI 12.6%-20.5%, I2 = 82.0%) and 33.3% at 3 years (95% CI 20.1%-49.8%, I2 = 0%). Early re-amputation after MLLA occurred in 10.9% of patients (95% CI 4.5%-24.4%, I2 = 40.3%). Baseline comorbidity burden was high in both groups, with substantial heterogeneity across studies. Conclusions: CLTI carries a poor prognosis regardless of treatment strategy. Infrainguinal bypass is associated with lower early mortality and better early limb preservation than primary MLLA, but long-term survival remains poor and later limb failure is common. Primary MLLA is not a low-risk alternative. Better contemporary comparative evidence utilising modern causal inference approaches is needed to support individualised decision-making.
Ricarte Almeida, E. R.; Mata Quintero, C. J.; Sesma Chazaro, J.; Peralta Rivera, C.; Arteaga Gonzalez, C. D.
Show abstract
Background: Sleeve gastrectomy is the most frequently performed bariatric procedure worldwide but is associated with the development of de novo gastroesophageal reflux disease (GERD). Hiatal hernia has been identified as a relevant anatomical factor in postoperative reflux, although most studies evaluate it dichotomously without analyzing whether its size influences GERD risk. The aim was to evaluate the association between preoperative hiatal hernia size and de novo GERD after sleeve gastrectomy. Methods: Retrospective, single - center, observational study of patients undergoing sleeve gastrectomy at Hospital Central Norte de Petroleos Mexicanos (2018 - 2025). Demographic and clinical characteristics, endoscopic classification of hiatal hernia size (small <2 cm, medium 2.1 - 4 cm, large >4 cm), and evidence of de novo GERD were analyzed using descriptive statistics, Fisher's exact test, odds ratio (=R) estimation with 95% confidence intervals (CI), and binary logistic regression. Statistical significance was set at p<0.05. Results: Fiftysix patients were included (mean age 48.3 {+/-} 8.1 years; 67.9% male). Hiatal hernia classification was conclusive in 46 patients (82.1%): 63.0% no hernia, 4.3% small, 30.4% medium, and 2.2% large. De novo GERD occurred in 14.0% of patients without preexisting GERD (6/43). No significant association was found between hiatal hernia size and de novo GERD (Fisher p=0.515). In the reduced logistic model, neither hiatal hernia (medium/large vs. absent/small; OR 3.47; 95% CI 0.50 - 29.43; p=0.207) nor age (OR 1.02; 95% CI 0.90 - 1.13; p=0.754) was significantly associated. No evaluated factor (sex, smoking, alcohol, age) reached significance. Conclusions: In this cohort, no statistically significant association was demonstrated between preoperative hiatal hernia size and de novo GERD after sleeve gastrectomy; however, the low number of events limits the ability to exclude a clinically relevant association. These findings are compatible with a multifactorial mechanism rather than with the isolated presence of this finding. Prospective studies with larger sample sizes and standardized reflux assessment instruments are required to confirm these results.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Chau, G. N.; Biswas, B. A.; Wagle, B. R.; Maeder, M. E.; Yu, J. B.; Bhattacharya, I.
Show abstract
Automated lesion segmentation is increasingly central to PSMA PET/CT interpretation, supporting staging, treatment planning, and response assessment at a scale that outpaces available nuclear-medicine expertise. However, automated PSMA-PET/CT whole-body lesion segmentation models are trained on images alone, with no knowledge of where in the body prostate metastases actually tend to occur. Radiologists use clinical domain knowledge of metastatic spread, but its absence in machine learning models produces false positives in anatomically implausible locations and missed lesions in high-risk sites such as the liver. In this work, we explore whether population-level spatial knowledge of metastatic spread can be used to augment deep learning segmentation predictions, and how such a prior should be fused with a network's output, without additional training. We build a data-driven metastasis atlas from 375 expert-annotated whole-body PSMA PET/CT scans and investigate its fusion with a trained segmentation network under a Bayesian framework, in which prediction probabilities from an nnU-Net-based lesion segmentation model serve as the likelihood and the data-driven atlas as the prior. Because metastases occupy only a small fraction of whole-body voxels, the atlas's peak probability is too low, and standard power-scaled or naive Bayesian pooling references lack the tools to deal with this shortcoming. This causes these standard fusion strategies to fail and, in the naive Bayesian case, to sharply degrade performance. We instead derive a calibrated, background-referenced log-odds fusion, one of many possible approaches to combine a population atlas with a deep learning model's predictions, distinct from classical multi-atlas label fusion in that it fuses a single population prior with a trained network's softmax rather than combining several registered atlases. Furthermore, this approach is neutral outside atlas support by construction, reduces exactly to the baseline network when unweighted, and requires no retraining. This atlas fusion significantly improved mean Dice over the baseline nnU-Net on a disjoint internal test set ($+0.011$, Holm-adjusted $p=0.021$) and on an independent external cohort ($+0.0129$, Holm-adjusted $p=3.8\times10^{-16}$), with lesion sensitivity improving from 0.849 to 0.861 internally and Dice improving over baseline in every stratified anatomic region, including the rare, high-risk sites motivating this work, while naive Bayesian pooling degrades performance sharply and power-scaled pooling underperforms it throughout. Our findings suggest that population-level spatial priors can meaningfully augment deep learning predictions in whole-body oncologic segmentation, provided the fusion rule is calibrated to where the prior actually carries signal.
Shi, Z.; Budhkar, A.; Amin, W.; Pollok, K. E.; Su, J.; Huang, K.
Show abstract
Improvements in data availability, sharing, and integration, together with the development of explainable artificial intelligence (XAI) techniques, are advancing precision medicine for pediatric cancer by facilitating diagnosis, biomarker discovery, and drug development. Data sharing commons and initiatives like the Childhood Cancer Data Initiative (CCDI) provide access to pediatric-specific genomic and clinical data cohorts and improve data availability for pediatric cancer research. Based on CCDI, a scalable AI platform, Graph Artificial Intelligence for Pediatric Oncology (GAIPO), integrates various data modalities from bulk and single-cell omics data to clinical information. Such multi-modal data facilitates the training and development of advanced XAI models for pediatric cancers. We then developed an end-to-end multi-modality framework, PCGS, for pediatric cancer by incorporating omics-specific representation learning via GNN models with cross-attention fusion and multi-objective learning for downstream tasks such as classification, clustering, and survival analysis. This framework outperforms previous supervised multi-omics integration baseline approaches based on glioma and Wilms tumor cohorts and enables GNN model explainability via Shapley value-based feature attribution approaches to explain the contributions of gene-level features across various biomedical tasks, including classification and survival. Given specific background samples (e.g., age groups, sex, grades) as baselines, this explainable GNN model estimates and ranks the importance scores for input features from each omics modality. It identifies background-specific key features for biomarker discovery, risk group identification, and survival analysis in glioma and Wilms tumor, with potential applicability to other pediatric cancers.
Gao, Y.; Yu, S.; Xia, Y.; Chen, S.; Xia, S.; An, R.; Zeng, J.; Zhao, F.; Ma, Y.; Wang, Y.; Xie, X.; Zhang, J.
Show abstract
Prognostic models in oncology are developed one cancer at a time, from that cancer's own labelled outcomes, and fail where prognostic information is scarcest. Rare cancers account for roughly a fifth of diagnoses and most paediatric malignancies, yet seldom supply enough events for a reliable time-to-event model. We therefore asked whether a representation learned without outcome labels can supply what those cohorts cannot. A Transformer encoder was pretrained by masked field-value modelling on 9425135 tumour records from the SEER 17 registries, diagnosed in 2000 to 2023. Only diagnosis-time fields passing a fail-closed coding-verification gate were admitted, and each record was emitted as an era-specific and a harmonised view, keeping two decades of recoding auditable. The encoder was then frozen and read by a linear Cox head for overall survival. Nine rare cancers were removed from the pretraining corpus entirely, each requiring an independent pretraining run. On a sealed test partition, all nine exceeded an architecture-identical random frozen encoder in Harrell concordance by +0.0034 to +0.0368, every lower confidence limit above zero. At 256 labelled patients, all 67 cancers favoured the pretrained representation over budget-matched Cox regression, median difference +0.0283. The advantage was bounded: given the entire training set, Cox regression was favoured in seven of nine rare cancers. The encoder did not outperform a field-frequency baseline on its own objective, so upstream reconstruction did not predict downstream transfer. Outcome-agnostic registry pretraining carries prognostic signal into cancers it has never seen, and is most useful where labels are fewest, without establishing clinical utility.
Zhao, L.; Zeng, Y.; Abelman, D. D.; Lin, W.; Luo, P.
Show abstract
Motivation: Cell-free DNA methylation provides a minimally invasive signal for early cancer detection and tissue-of-origin prediction. Most methods represent methylation measurements as independent fixed-window features and therefore do not explicitly model relationships among genomic regions. Results: We developed PANGEM (Pan-cancer Graph-based Cancer Detection Using the Cell-free DNA Methylome), a graph-learning framework that represents genomic bins as nodes and integrates CpG context, genomic proximity, and sample-specific methylation similarity in the graph topology. Across five repeated stratified train-test splits, PANGEM achieved the highest mean performance among evaluated methods, with an AUROC/AUPR of 0.997/1.000 for binary cancer detection and macro-AUROC/AUPR of 0.977/0.870 for multiclass tissue-of-origin prediction. In the independent INSPIRE cohort, 72 of 78 cancer cases (92.3%) exceeded the binary classification threshold, and PANGEM correctly classified 9 of 17 head and neck cancer cases (52.9%), the highest accuracy among evaluated methods. Subnetwork analysis further identified recurrent, graph-connected methylation patterns, including a 111-DMR subnetwork with increased methylation in cancer samples.
dela Sotta, T.; Saavedra, J. M.; Chang, V.; Xavier, A.; Henriquez, H.; Orellana, Y.; Curimil, J.
Show abstract
Diffusion models achieve high reconstruction quality in low-dose computed tomography (LDCT), but their iterative sampling trajectories impose substantial computational costs. Unlike unconditional generation, paired LDCT reconstruction starts from an image that already contains the anatomy and spatial structure of the standard-dose CT (SDCT) target; reconstruction primarily requires correcting dose-related noise and artifacts. We therefore introduce Residual Endpoint Flow Matching (REFM), an LDCT reconstruction method that learns to transport an LDCT image directly toward its paired SDCT endpoint rather than defining a noise-to-image trajectory. REFM predicts the residual velocity along linear interpolations between both images and supports single-step and multi-step reconstruction using the same trained network. We evaluate five model capacities using 1 to 50 Euler steps against deterministic U-Net and diffusion-based baselines. Across all REFM variants, one-step inference consistently provides the highest reconstruction quality. On the TCIA validation set, REFM Base achieves 50.98 dB PSNR and 0.9865 SSIM at 94.54 fps, compared with 50.92 dB, 0.9847, and 9.26 fps for DDPM-10. REFM Small retains 50.71 dB while increasing throughput to 198.56 fps. Without fine-tuning, REFM Base also matches the 25-step DDPM baseline on the external Mayo Clinic dataset, although DDPM remains stronger on synthetically degraded CRLM images. Thus, our results show that exploiting paired anatomical correspondence enables diffusion-level LDCT reconstruction with a single step reconstruction.
Corcoran, D.; Szoeke, C.; Apostolopoulos, V.; Feehan, J.
Show abstract
This study aimed to quantify the longitudinal tracking and cross-sectional construct validity of a single-item questionnaire measuring recreational physical activity frequency (RPAF) in the Womens Healthy Ageing Project. At baseline, 474 participants aged 45-55 reported RPAF from 1993 to 2014. Longitudinal tracking of the RPAF item was assessed as a consecutive-wave and baseline-referenced measure using linear weighted kappa (LWK), Spearman correlations, exact agreement and within-one-category agreement. Construct validity in the form of convergent and known-group validity was assessed using the International Physical Activity Questionnaire (IPAQ) leisure activity domains, Short Form 36 physical function (SF-36-PF) subscale, Timed Up and Go (TUG), hand grip strength (HGS) and waist-to-height ratio (WHtR). 474 participants provided baseline RPAF data. Pairwise longitudinal samples ranged from 176 to 459 across the study. Consecutive-wave LWK ranged from 0.38 to 0.49, and Spearman correlations ranged from 0.44 to 0.57. Exact and within-category agreement ranged from 41.4%-50.8% and 72.0%-79.0%. Baseline-referenced LWK ranged from 0.22 to 0.47, with Spearman correlations of 0.29 to 0.56. RPAF correlated with total IPAQ leisure score (rs = 0.60), IPAQ walking score (rs = 0.58), SF-36-PF (rs = 0.33) and TUG score (rs = -0.25). No significant correlation was identified between RPAF, HGS or WhTR. RPAF discriminated known groups for WHO guideline-sufficient activity, SF-36-PF, and TUG fall risk. The RPAF item demonstrated fair-to-moderate agreement in consecutive waves, with weaker baseline-referenced tracking. Cross-sectional validity was highest with total IPAQ leisure activity. The item may provide a pragmatic measure for RPAF in womens cohort studies.
Humphries, C.; Brett, J.; Gruber, F.; James, E.; McKendrick, T. I.; McNairn, K. C.; Miell, A.; O'Brien, R.; Rahman, F.; Schölin, L.; Stewart, M.; Casey, A.
Show abstract
Objective To measure the accuracy of clinical coding, clinician review, and a locally deployed large language model (LLM) in identifying alcohol, drug, and self-harm involvement in emergency department (ED) attendances, and quantify prevalence. Design Two-phase diagnostic accuracy study. In a validation week, the identification strategies were assessed against a conflict-adjudicated reference standard (n=2,256); the LLM was then applied to n=105,096 annual attendances at the same site. Setting UK Type 1 Emergency Department treating patients [≥]16yrs. Main outcome measures Prevalence quantification compared with the reference standard; sensitivity, specificity, and balanced accuracy of each strategy; monthly identification rates and adjusted annual prevalence. Results The reference standard identified 12.1% of attendances as involving alcohol, drugs, or self-harm (coding 6.0%; clinician 10.0%, LLM 15.6%). LLM balanced accuracy matched or outperformed clinician review in all three domains (alcohol 0.942 v 0.930, p=0.635; drug 0.959 v 0.791, p<0.001; self-harm 0.982 v 0.908, p=0.004). Coding recorded 1.07 domains per identified patient against 1.32 in the reference standard. Adjusted annual prevalence corresponded to 12,890 domain involvements per year not identifiable in coded data. Subdomain classification found at least 81.6% of self-harm attendances required medical assessment for injury or overdose before psychiatric review. Conclusions Clinical coding identified fewer than half of presentations involving alcohol, drugs, and self-harm and rarely captured co-occurring domains; under-recording was present across a full year. A locally deployed LLM generated more complete structured data from existing clinical text within NHS infrastructure, at a scale which is not feasible for manual review.
Yao, R.; Wi, C.-I.; Beenken, M. J.; Watson, D.; Wheeler, P. H.; Finch, M.; Kelleher, D. P.; Anil, G.; Anderson, T.; Madden, K.; Okuno, S. H.; Odedina, F. T.; Westfall, E. C.; Park, E. Y.; Sharma, P.; Dugani, S.; Foss, R. M.; Hidaka, B. H.; Sosso, J. L.; Sabarish, S.; Singh, G.; Lugo-Fagundo, N.; Howick, J.; Kim, W. R.; Calvin, A. D.; Walker-Mcgill, C. L.; Rennert, L.; Juhn, Y. J.; Cerhan, J. R.; Lynch, B. A.
Show abstract
Purpose: This study assesses the association between colorectal cancer (CRC) screening and a validated, housing-based measure of individual-level socioeconomic status (SES, called HOUSES hereafter) within rural communities and determines whether HOUSES-integrated geospatial analysis can be used to tailor interventions. Methods: We used CRC screening data from a subset of Mayo Clinic Midwest patients living in cities without ready access to routine care in the Mayo Clinic Health System in 2019 to represent rural communities. At the individual level, we assessed the association between CRC screening rates and the HOUSES index, adjusting for age, sex, race/ethnicity, comorbidity, distance from home address to clinic, and area deprivation index, using a multilevel mixed-effects logistic regression model. Additionally, we conducted geospatial analysis to examine the correlation between hotspots of 1) lower CRC screening rates and 2) lower SES of the subject population (HOUSES quartile 1). Findings: Among 34,489 individuals (median age 64.0 years, 52.4% female), those with the lowest SES (HOUSES Q1) had 37% lower odds of being CRC screening adherent than those with the highest SES (HOUSES Q4) (adj. OR [95% CI]: 0.63 [0.58-0.69]). In the 14 identified HOUSES Q1 hotspots, there was a significant correlation in counts of HOUSES Q1 and low CRC screening (correlation coefficient=0.81). Conclusion: Lower SES was significantly associated with lower CRC screening among rural populations. HOUSES-enabled geospatial analysis identified geographic hotspots with lower CRC screening rates for targeted interventions to address disparities in CRC screening in rural communities. HOUSES may be a useful digital tool for cancer preventive care and research.
De Luca, S.; Fava, C.; Rizzo, G.; Visconti, A.; Berchialla, P.
Show abstract
Background. Patient stratification from multi-omics and clinical data is essential for uncovering disease heterogeneity and moving toward more personalized treatment strategies. However, integrating heterogeneous data layers while identifying robust patient strata remains challenging. Methods. We introduce Reduced Fusion of Multi-Omics Stratification (RedFuMOS), a novel three-step approach for patient stratification based on mixed-type multi-omics data. RedFuMOS extends Similarity Network Fusion to accommodate mixed-type data layers and layer-specific similarity measures for data integration, includes a dimensionality reduction step to mitigate the curse of dimensionality, and performs patient stratification using density-based hierarchical clustering with HDBSCAN. It also implemented an automated optimization procedure to identify the best set of hyperparameters, minimizing the need for manual tuning. Results. RedFuMOS outperformed six state-of-the-art tools for multi-omics patient stratification in a comprehensive simulated benchmarking study, which also confirmed that, although computationally expensive, the dimensionality reduction step is crucial for achieving good stratification performance. Additionally, RedFuMOS identified two clinically relevant patient strata in a small real-world cohort of patients with Philadelphia chromosome-positive chronic myeloid leukaemia. Conclusion. RedFuMOS provides a flexible framework for integrating heterogeneous multi-omics and clinical data. RedFuMOS is available as an R package at http://github.com/delucasara/RedFuMOS.
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.